Papers with large-scale conceptualization taxonomies
MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation Dataset (2025.acl-long)
Copied to clipboard
| Challenge: | Recent advances in LLMs have demonstrated superior performance in a variety of reasoning tasks (Liu et al., 2023b; Chan e t al, 2024; Qin eetal., 2023) However, to truly achieve conscious processing, the integration of System II reasoning ability is essential. |
| Approach: | They propose a three-step process for reasoning with distributional changes, termed as a metaphysical resoning, and propose 'MARS' task to assess LLMs' reasoning abilities. |
| Outcome: | The proposed task is based on a three-step discriminative process and is compared with a standard model with 20 LLMs of varying sizes and methods. |